Papers with Automated Essay Scoring

22 papers
Beyond the Gold Standard in Analytic Automated Essay Scoring (2025.acl-srw)

Copied to clipboard

Challenge: Automated Essay Scoring (AES) is a new approach to assessing writing practice . traditional holistic scoring methods are not reliable and lack formative feedback in the classroom.
Approach: They propose to combine analytic and holistic AES to create a system that learns from individual raters instead of gold standard labels.
Outcome: The proposed system learns from individual raters instead of gold standard labels.
Multi-task Learning for Automated Essay Scoring with Sentiment Analysis (2020.aacl-srw)

Copied to clipboard

Challenge: Automated Essay Scoring (AES) is a process that aims to alleviate the workload of graders and improve the feedback cycle in educational systems.
Approach: They propose to combine two tasks, sentiment analysis and AES by utilizing multi-task learning to combine sentiment features extracted from opinion expressions.
Outcome: The proposed model produces a QWK of 0.763 on the Automated StudentAssessment Prize (ASAP) benchmark.
Neural Automated Essay Scoring and Coherence Modeling for Adversarially Crafted Input (N18-1)

Copied to clipboard

Challenge: Existing approaches to Automated Essay Scoring (AES) are not well-suited to capture adversarially crafted input of grammatical but incoherent sequences of sentences.
Approach: They propose a neural model of local coherence that can effectively learn connectedness features between sentences.
Outcome: The proposed approach strengthens the validity of neural essay scoring models.
Qayyem: A Real-time Platform for Scoring Proficiency of Arabic Essays (2026.acl-demo)

Copied to clipboard

Challenge: Existing Arabic writing technologies primarily use a single quality score for essays, but there is limited support for Arabic AES.
Approach: They propose a Web-based platform that integrates Arabic AES workflows with a user-friendly interface.
Outcome: The proposed system integrates with existing Arabic scoring systems and provides a user-friendly interface.
Enhancing Marker Scoring Accuracy through Ordinal Confidence Modelling in Educational Assessments (2025.acl-industry)

Copied to clipboard

Challenge: Automated Essay Scoring (AES) systems aim to evaluate the quality of candidate writing using computational methods.
Approach: They propose a model that assigns a confidence score to each automated score to ensure it meets high reliability standards.
Outcome: The proposed model achieves an F1 score of 0.97 and releases 47% of predicted scores with 100% CEFR agreement and 99% with at least 95% CEFR agreeance compared to the standalone model where all predicted scores are released.
Enhancing Automated Essay Scoring Performance via Fine-tuning Pre-trained Language Models with Combination of Regression and Ranking (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent work on sentence prediction tasks uses shallow neural networks to learn essay representations and constrain calculated scores with regression loss or ranking loss.
Approach: They propose to use a pre-trained language model to learn text representations first and then to constrain the scores with regression loss or ranking loss.
Outcome: The proposed model outperforms state-of-the-art models on the Automated Student Assessment Prize dataset.
LAILA: A Large Trait-Based Dataset for Arabic Automated Essay Scoring (2026.eacl-long)

Copied to clipboard

Challenge: Existing Arabic resources are small in scale and lack trait-specific annotations.
Approach: They propose to use LAILA to build a large Arabic AES dataset with holistic and trait-specific annotations of seven writing proficiency traits.
Outcome: The LAILA dataset comprises 7,859 essays annotated with holistic and trait-specific scores on seven dimensions: relevance, organization, vocabulary, style, development, mechanics, and grammar.
TCFLE-8: a Corpus of Learner Written Productions for French as a Foreign Language and its Application to Automated Essay Scoring (2023.emnlp-main)

Copied to clipboard

Challenge: Automated Essay Scoring (AES) aims to automatically assess the quality of essays.
Approach: They propose to use a corpus of 6.5k essays collected in the context of the Test de Connaissance du Français (TCF) certification exam to foster the development of AES for French.
Outcome: The proposed system can assess the quality of essays in a language certification exam using a corpus of 6.5k essays collected in the TCFLE-8 exam.
On the Use of Bert for Automated Essay Scoring: Joint Learning of Multi-Scale Essay Representation (2022.naacl-main)

Copied to clipboard

Challenge: Pre-trained models have not been used to outperform other deep learning models such as CNN in Automated Essay Scoring (AES).
Approach: They propose a novel multi-scale essay representation for BERT that can be jointly learned . they employ multiple losses and transfer learning from out-of-domain essays to further improve performance .
Outcome: The proposed model outperforms existing models in the area of automated essay scoring . the proposed model generalizes well to the CommonLit Readability Prize data set .
Representation-to-Creativity (R2C): Automated Holistic Scoring Model for Essay Creativity (2025.findings-naacl)

Copied to clipboard

Challenge: Existing studies on Automated Essay Scoring (AES) are limited.
Approach: They propose a new essay rubric specifically designed for assessing creativity in essays . they use a ground truth data set to construct a self-supervised learning model .
Outcome: The proposed model improves the assessment of creativity in essays by 58% compared to the current models.
EssayJudge: A Multi-Granular Benchmark for Assessing Automated Essay Scoring Capabilities of Multimodal Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Automated Essay Scoring (AES) systems face three major challenges: reliance on handcrafted features that limit generalizability, difficulty in capturing fine-grained traits like coherence and argumentation, and inability to handle multimodal contexts.
Approach: They propose a multimodal benchmark to evaluate AES capabilities across lexical-, sentence-, and discourse-level traits without manual feature engineering.
Outcome: The proposed system can evaluate AES capabilities across lexical-, sentence-, and discourse-level traits without manual feature engineering.
PsyScore: A Psychometrically-Aware Framework for Trait-Adaptive Essay Scoring and ZPD-Scaffolded Feedback (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to Automated Essay Scoring (AES) treat scoring and feedback as separate components, resulting in fragmentation.
Approach: They propose a psychometrically-aware framework that integrates diagnostic assessment with instructional scaffolding through a shared latent ability representation.
Outcome: The proposed framework integrates diagnostic assessment with instructional scaffolding through a shared latent ability representation.
Computing with Subjectivity Lexicons (2020.lrec-1)

Copied to clipboard

Challenge: a new set of lexicons for expressing subjectivity in text documents is presented . lexiconics are useful resources for identifying semantics relevant to sentiment, emotion, personality, language bias, mood, and attitude.
Approach: They propose a set of lexicons for expressing subjectivity in Brazilian Portuguese text documents . they use word embedding techniques to capture semantically related words to the ones in the lexicos .
Outcome: The proposed lexicons represent different subjectivity dimensions and are more compact in number of terms.
Beyond Agreement: Diagnosing the Rationale Alignment of Automated Essay Scoring Methods based on Linguistically-informed Counterfactuals (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing Automated Essay Scoring (AES) methods focus on sentence-level features, whereas Large Language Models (LLMs) are sensitive to conventions & accuracy, language complexity, and organization.
Approach: They propose to use large language models to aid in decision-making . they propose to analyze the reasoning of neural models by analyzing sentence-level features.
Outcome: The proposed method improves understanding of neural approaches to Automated Essay Scoring (AES) and can also apply to other domains seeking transparency in model-driven decisions.
Improving Domain Generalization for Prompt-Aware Essay Scoring via Disentangled Representation Learning (2023.acl-long)

Copied to clipboard

Challenge: Existing AES models are either prompt-specific or prompt-adaptive and cannot generalize well on “unseen” prompts.
Approach: They propose a prompt-aware neural AES model to extract comprehensive representation for essay scoring, including both prompt-invariant and prompt-specific features.
Outcome: The proposed model extracts comprehensive representation for essay scoring, including both prompt-invariant and prompt-specific features.
Aggregating Multiple Heuristic Signals as Supervision for Unsupervised Automated Essay Scoring (2023.acl-long)

Copied to clipboard

Challenge: Automated Essay Scoring (AES) aims to evaluate the quality score of input essays without human intervention.
Approach: They propose an unsupervised approach to evaluate the quality of input essays . they use multiple heuristic quality signals as pseudo-groundtruths to train a neural AES model .
Outcome: The proposed approach achieves state-of-the-art performance on eight prompts of ASPA dataset compared with previous unsupervised methods .
Mixture of Ordered Scoring Experts for Cross-prompt Essay Trait Scoring (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches to automate essay scoring overlook critical information, authors say . evaluators often limit their performance to unseen topics, resulting in incomplete assessment perspectives.
Approach: They propose a framework that integrates information from prompts and essays into an AES framework.
Outcome: The proposed framework achieves state-of-the-art in cross-prompt scoring and multi-trait scoring on the ASAP++ dataset.
Towards Explainable Chinese Native Learner Essay Fluency Assessment: Dataset, Tasks, and Method (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing GEC datasets in Chinese fail to consider specific grammatical error types and overlook cross-sentence grammamatical errors.
Approach: They propose to use Chinese essay fluency assessment to assess essay fluencies along with coarse and fine-grained errors and corrections to improve explainability.
Outcome: The proposed dataset encapsulates essay fluency scores along with both coarse and fine-grained errors and corrections.
Beyond the Score: Uncertainty-Calibrated LLMs for Automated Essay Assessment (2025.emnlp-main)

Copied to clipboard

Challenge: Automated Essay Scoring (AES) systems attain near–human agreement on some public benchmarks, but real-world adoption is limited.
Approach: They propose a distribution-free wrapper that equips any classifier with set-valued outputs enjoying formal coverage guarantees.
Outcome: The proposed model achieves coverage targets while keeping prediction sets compact.
MAPLE: A Meta-learning Framework for Cross-Prompt Essay Scoring (2026.findings-acl)

Copied to clipboard

Challenge: Current approaches to automate essay scoring (AES) treat each writing task as a separate task, resulting in inconsistent performance.
Approach: They propose a meta-learning framework that leverages prototypical networks to learn transferable representations across different writing prompts.
Outcome: The proposed framework outperforms baseline models on ELLIPSE and ASAP (English) and LAILA (Arabic) on three diverse datasets.
Transformer-based Joint Modelling for Automatic Essay Scoring and Off-Topic Detection (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies show that automated essay scoring systems assign lower grades to irrelevant responses.
Approach: They propose an unsupervised technique that jointly scores essays and detects off-topic essays.
Outcome: The proposed method outperforms baseline and earlier conventional methods on two essay-scoring datasets in off-topic detection and on-topic scoring.
Graph-Based Multi-Trait Essay Scoring (2025.emnlp-main)

Copied to clipboard

Challenge: Existing work on Automated Essay Scoring (AES) models essay as word sequence, but new approach uses graph-attention network approach to model essay traits.
Approach: They propose a graph-attention network approach to automate essay scoring that models interactions among essay traits as a graphical graph.
Outcome: The proposed approach outperforms competing approaches on the ASAP++ dataset . it allows for multiple-task scoring, allowing for more detailed feedback on essays .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations